Papers with problem-solving process

8 papers
A Comprehensive Evaluation of Inductive Reasoning Capabilities and Problem Solving in Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Inductive reasoning is fundamental to both human and artificial intelligence.
Approach: They evaluated the inductive reasoning abilities of current Large Language Models (LLMs) and their performance on symbolic tasks.
Outcome: The proposed models fail on symbolic tasks and show that chain-of-thought prompts help them by decomposing the problem-solving process, but the LLMs learn limitedly.
Advancing Reasoning with Off-the-Shelf LLMs: A Semantic Structure Perspective (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing reasoning models suffer from hallucinations and unfaithfulness, whereas general LLMs perform suboptimal on complex tasks.
Approach: They propose a structure analysis method that helps LLMs better understand the question structure and guide the problem-solving process.
Outcome: The proposed method improves zero-shot performance on knowledge-intensive and mathematical tasks while demonstrating strong robustness against corrupted reasoning paths.
MuMath: Multi-perspective Data Augmentation for Mathematical Reasoning in Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) that integrate with external Python interpreters are not able to demonstrate the calculation process, which compromises user-friendliness and understanding of problem-solving steps.
Approach: They propose to use LLaMA-2 to refine LLti-perspective augmentation methods to improve performance.
Outcome: The proposed model achieves 88.3% on GSM8K and 34.5% on MATH.
Scaling Collaborative Effort with Agents (2026.findings-acl)

Copied to clipboard

Challenge: Current evaluations of agents focus on producing high-quality, final outputs in one shot, failing to account for the inherently iterative nature of many real-world problems.
Approach: They propose a framework that captures how an agent’s utility grows with increasing user involvement.
Outcome: The proposed framework captures how an agent’s utility grows with increasing user involvement, revealing a missing ingredient in agent design: the ability to sustain engagement and scaffold user understanding.
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances on prompting and post-training have enabled LLMs to perform step-wise reasoning tasks, but they tend to explore unproductive solution paths without effective backtracking or strategy adjustment.
Approach: They propose a framework that empowers LLMs to “think about how to think” and dynamically adapts reasoning strategies in real-time.
Outcome: The proposed framework outperforms previous SOTA methods by 9-12% in accuracy while reducing inference time by 28-35% under the same compute budget.
Reasoning Graph Enhanced Exemplars Retrieval for In-Context Learning (2025.coling-main)

Copied to clipboard

Challenge: Existing methods focus on semantic similarity between queries and candidate exemplars, while logical connections between reasoning steps can be beneficial to depict problem-solving process.
Approach: They propose a method to retrieve exemplars with semantic and structural similarity using a graph kernel.
Outcome: The proposed method is superior to state-of-the-art retrieval-based approaches on mathematics and logical reasoning tasks.
Pedagogical Alignment of Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often used without pedagogical fine-tuning and provide immediate answers rather than guiding students through the problem-solving process.
Approach: They propose a method for constructing large-scale preference datasets using synthetic data generation techniques that eliminates the need for manual annotation.
Outcome: The proposed methods outperform standard supervised fine-tuning (SFT) and improve alignment accuracy by 13.1% and 8.7% respectively.
SedarEval: Automated Evaluation using Self-Adaptive Rubrics (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation paradigms rely on generic scoring rubrics that fail to consider the specificities of each question and its problem-solving process.
Approach: They propose a new evaluation paradigm based on self-adaptive rubrics that mimic a human evaluator's analytical process.
Outcome: The proposed evaluation paradigm achieves higher concordance rate with human graders than existing paradigms, including GPT-4.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations